Frontiers in Artificial Intelligence
○ Frontiers Media SA
Preprints posted in the last 7 days, ranked by how well they match Frontiers in Artificial Intelligence's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Oyarzun Silva, R.; Hernandez Hernandez, P.
Show abstract
Background. Accurate delineation of the gross tumour volume (GTV) - primary tumour (GTVp) and nodal disease (GTVn) - on FDG-PET/CT is a critical step of head and neck radiotherapy planning. Comparisons between lightweight custom networks and the auto-configured nnU-Net v2 are usually reported as end-to-end pipelines, conflating the contribution of the network with that of the inference-time post-processing applied on top of it. We separated the two. Methods. MiniUNet3D (custom 3D U-Net, 18.3 M parameters) and nnU-Net v2 (3d_fullres, 88.2 M parameters) were trained on the same 578 FDG-PET/CT cases (85/15 author-defined split of the HECKTOR 2025 Task 1 set, 8 centres) and evaluated on the same internal cohort. Three arms were compared pairwise: MiniUNet3D raw output at a fixed 0.5 threshold, MiniUNet3D with a locked adaptive post-processing pipeline, and nnU-Net v2. Comparisons used paired Wilcoxon tests with bootstrap confidence intervals, Bonferroni and Benjamini-Hochberg correction, and Cohen's d; catastrophic failure (Dice < 0.01) was compared with an exact McNemar test. Cases with an empty reference for a given target were excluded from that target's analysis (n = 98 GTVp, n = 93 GTVn). Results. With post-processing matched off, nnU-Net v2 was superior: median GTVp Dice 0.799 versus 0.592 (mean difference -0.244, 95 % CI -0.300 to -0.191; d = -0.88) and GTVn 0.774 versus 0.598 (d = -0.82). Post-processing raised MiniUNet3D to 0.800 (GTVp) and 0.738 (GTVn), recovering 79 % of that difference. Post-processed, MiniUNet3D matched nnU-Net v2 on GTVp Dice (p = 0.113) but remained inferior on nodal disease after Bonferroni correction (Dice p = 0.041; surface Dice p = 0.049). Catastrophic GTVp failures were 25/98 raw, 8/98 post-processed and 1/98 for nnU-Net v2 (McNemar p = 0.016). Inference took 34 s versus 78 s per case on the same GPU. Conclusions. Post-processing recovered most, but not all, of the difference between the two models, and it did not confer robustness: an eight-fold higher rate of empty contours on small primaries persisted, which is the more consequential difference for planning safety. Pipeline comparisons reported without a post-processing ablation risk attributing to a network what post-processing supplied.
bolin, k.; Stibrant Sunnerhagen, K.
Show abstract
Background The time trend in long-term survival after a stroke is to some extent unknow due to (relatively) short follow up periods in available data. The objective of this study is to identify and quantify differences in long-term stroke survival in Sweden between men and women and patients with different attained educational levels, comparing two time-periods, 2000-2009 and 2010-2022. Methods This study employs total population Swedish register data pertaining to hospital-based care and mortality due to stroke for the period 2000-2022 in order to estimate survival (all-cause mortality) after ischaemic and haemorrhagic stroke, respectively, and pertaining to attained educational level. Kaplan-Meier survival functions are estimated stratifying for time-period, sex and educational level. Cox regressions are employed to quantify mortality hazard ratios between the strata. Age is taken into account in complementary analyses (supplement). Results Taking only time-period (2000-2009 vs 2010-2022) into account resulted in significantly higher survival in the second period for ischaemic stroke patients (HR: 0.84; 95% CI: 0.83-0.84), while no significant difference could be detected for haemorrhagic stroke. Stratifying for sex showed that men gained more than women in terms of reduced mortality hazard rate between the periods. Further stratifying by educational level and estimating survival separately for men and women showed that, for both men and women, patients with the lowest education were relatively worse off (compared to patients with higher education) in the second period. Further analyses, taking age into account, reversed the relative hazard ratio between men and women, but corroborated the result that low education is associated with poorer outcome than high education. Conclusions The results suggest that there are considerable differences in expected long-term survival after stroke between the sexes, but that this may be due to differences in age between the sexes at the time of stroke. Moreover, lower educational level is significantly associated with lower long-time survival.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Sadia, H.; Doyon, N.; Duchesne, S.
Show abstract
Background Understanding the mechanisms underlying brain aging and age-related pathological changes is essential for advancing brain health research. Our group previously developed a mechanistic mathematical model of healthy brain, Chamberland et al. (2024) that integrates key biological processes involved in normal aging, from which Alzheimer's disease (AD) related changes may emerge naturally. Objectives To characterize and validate this brain model by evaluating its sensitivity, calibrating its parameters, and assessing generalizability in independent populations. Methods The model represents the evolution of key biological processes associated with brain aging, including amyloid beta (A{beta}), tau pathologies, neuroinflammation, and neuronal death. After identifying the 30 most influential parameters, we calibrated the model using cognitively normal (CN) participants from the AD Neuroimaging Initiative (ADNI) database (n = 211) by minimizing a loss function composed of three outcomes (AB) plaques, tau tangles, and neuronal density). The calibrated model was then applied to the UK Biobank cohort (n = 35,899) of normal controls (aged 44-82 years). The effects of sex and APOE were evaluated using stratified simulations. Results Parameter calibration significantly reduced the prediction errors for A{beta} and tau. Neuronal density predictions showed strong agreement in the UK Biobank cohort. The variance decomposition identified APOE status as a major contributor to variability in A{beta}. Conclusion Our validated brain health model links mechanistic pathways with population data and reproduces neuronal density patterns in an independent cohort. These findings support its use as a framework for studying brain aging and investigating how Alzheimer's disease related pathological changes may emerge with aging.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.
Show abstract
Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.
Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O. R.; Klang, E.; Barash, Y.
Show abstract
Objective: Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it. Methods: We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses. Results: No prior's 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior's output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment). Conclusion: A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage. Significance: Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output's confidence identified high-risk selections better than the tested purpose-built ranker.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Alve, S. R.; Rahman, S.; Meem, S. M. A. C.
Show abstract
A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.
Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.
Show abstract
Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.
Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.
Show abstract
Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.
Show abstract
Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.
Show abstract
Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.